Skip to content

feat: Phase 4 — Boman scorer, consensus report, 89 AMP nominees, nomination report - #13

Merged
cschanhniem merged 5 commits into
mainfrom
feat/phase4-rebase
Jun 27, 2026
Merged

cschanhniem merged 5 commits into
mainfrom
feat/phase4-rebase

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

  • Boman index second scorer (src/openamp_foundry/scoring/boman.py): independent activity signal based on Boman (2003) interaction potentials, tanh-normalized to [0,1]; GRAVY score (Kyte & Doolittle 1982) added as physicochemical feature
  • Model disagreement signal: |activity_likeness − boman_activity| flags high-uncertainty candidates (≥0.30) and confirms dual-scorer consensus (<0.20)
  • Scorer consensus report (batch pack section 5): classifies each of the 89 nominees as high_consensus / moderate / uncertain
  • 89 AMP nominees produced by make phase3 with evidence certificates validated against schemas/candidate.schema.json
  • docs/NOMINATION_REPORT.md (441 lines): full scientific methodology — seed rationale, generation strategies, scoring formulas, selection criteria, reproducibility instructions, integrity section ("what we do not know"), safety assessment, reference list
  • docs/EXPERT_REVIEW_PACK.md updated with live Phase 3 data: top-20 table with Boman and disagreement columns, mean Boman=0.503, mean disagreement=0.311, suggested 5-candidate pilot

Commits (cherry-picked from feat/phase4-lab-bridge onto clean main)

  1. feat: Phase 4 lab-bridge — 45-AMP reference set, expert review pack, lab results schema, METHODS appendix
  2. feat: add Boman index second activity scorer and GRAVY feature — boman.py, physchem.py update, pipeline integration, test_boman_scorer.py (37 tests)
  3. feat: add scorer consensus report to batch pack (section 5) — batch_pack.py Section 5, test_batch_pack.py expanded
  4. docs: update EXPERT_REVIEW_PACK with 89-candidate Phase 3 nomination results
  5. docs: add NOMINATION_REPORT.md — full methodology behind 89 AMP nominees

Test plan

  • make test — all 314+ tests pass (ran on feat/phase4-rebase before push)
  • make phase3 — produces 89 selected candidates
  • python -m openamp_foundry.evidence.validate_certs --cert-dir outputs/certs — all certs validate against schema
  • Review docs/NOMINATION_REPORT.md for scientific accuracy
  • Review docs/EXPERT_REVIEW_PACK.md for completeness of top-20 table

Safety checklist

  • No wet-lab protocols included
  • No dangerous pathogen instructions
  • No toxicity-maximization objectives
  • All scoring is transparent heuristics with explicit disclaimers
  • Evidence certificates include known failure modes

🤖 Generated with Claude Code

…lab results schema, methods appendix

Builds the bridge between computational nomination (Phase 3) and wet-lab validation (Phase 4).

Changes:
  - examples/known_reference/amp_curated_references.csv: 45 diverse known AMPs from published
    literature (magainin, buforin, temporin, aurein, cecropin, indolicidin, cathelicidin families)
    for meaningful novelty scoring. Phase 3 re-run: mean novelty 0.139 → 0.172 vs real AMP space.
  - schemas/lab_result.schema.json: JSON schema for ingesting wet-lab assay results
    (MIC, MBC, hemolysis, cytotoxicity) — enables active-learning loop when data arrives.
  - src/openamp_foundry/data/lab_results.py: loader, validator, summariser, and candidate mapper
    for lab results ingestion.
  - docs/EXPERT_REVIEW_PACK.md: complete expert review pack ready to send to a qualified
    microbiologist. Includes batch stats, top-20 table, reviewer questions, next-step checklist.
  - docs/METHODS.md: publication-quality methods appendix covering generation, scoring,
    selection, reproducibility, benchmark validation, and known failure modes.
  - Makefile: phase3 now references amp_curated_references.csv instead of seeds.

15 new lab_results tests. 266 total tests pass. Lint clean.
make test && make demo both pass.

Remaining Phase 4 human gates (not automatable):
  - Expert review sign-off (docs/EXPERT_REVIEW_PACK.md)
  - CRO/lab partner selection
  - Synthesis and assay execution
  - Results ingestion via schemas/lab_result.schema.json
Implements the Boman (2003) interaction-potential-based activity scorer as
a second independent predictor alongside the existing physicochemical heuristic.
Adds model disagreement signal to flag uncertain nominations.

- scoring/boman.py: boman_index(), boman_activity_score(), gravy_score(), model_disagreement()
  - Published Boman 2003 Table 1 potentials (transparent, no training)
  - tanh normalization to [0,1]; disagreement = |activity − boman_activity|
- features/physchem.py: boman_index and gravy added to compute_features() output
- pipeline.py: boman_activity and disagreement stored in raw_scores
- scoring/ensemble.py: disagreement-aware selection reasons and failure modes
- schemas/candidate.schema.json: boman_activity and disagreement as optional score fields
- tests/test_boman_scorer.py: 37 tests covering all four functions + pipeline integration
- docs/METHODS.md, EXPERT_REVIEW_PACK.md: document second scorer and uncertainty signal

Candidates with disagreement < 0.20 have dual-scorer consensus (more robust).
Candidates with disagreement >= 0.30 are flagged for extra scrutiny.
Surfaces the Boman index vs. activity-likeness disagreement signal in the
human-readable batch pack and expert review documents.

- batch_pack.py: scorer_consensus_report() — new 5th sub-report
  - Labels each candidate: high_consensus (<0.20), moderate, uncertain (≥0.30)
  - Sorted by disagreement ascending (strongest consensus first)
  - Gracefully handles candidates without boman_activity in scores
  - generate_batch_pack() → batch_pack_version 1.1, includes scorer_consensus
  - Summary adds n_high_consensus, n_uncertain_disagreement, mean_scorer_disagreement
  - write_batch_pack_markdown() → Section 5 Scorer Consensus table
- tests/test_batch_pack.py: 11 new tests (TestScorerConsensusReport)
  - Updated _make_candidate helper to include boman_activity/disagreement
  - Updated TestWriteBatchPackMarkdown to assert "Scorer Consensus" in markdown
  - Updated TestGenerateBatchPack to assert scorer_consensus key present

314 tests pass.
…results

Reflects the actual Phase 3 run output: 89 candidates selected, all evidence
certificates schema-validated, dual-scorer consensus data from v0.2 pipeline.

Key changes:
- Top-20 table now includes Boman activity and disagreement columns (live data)
- Batch statistics include mean Boman (0.503), mean disagreement (0.311)
- Explains scientifically why high disagreement is expected for helical AMPs:
  activity scorer rewards amphipathic character; Boman index penalizes hydrophobic
  residues — these are different mechanistic models, not a data quality issue
- Adds "Suggested pilot candidates" table ranked by Boman × low-disagreement:
  SEED-003 tryptophan-rich 11-mers prioritized for first synthesis round
- Reviewer questions updated to cover dual-scorer methodology
- Limitations table updated to reflect current two-scorer state
- 6.2 asks expert: do SEED-003 11-mers look more promising than SEED-005 14-mers?
Complete scientific documentation of how the 89 Phase 3 candidates were
found, scored, filtered, and selected. Suitable for sharing with expert
reviewers and as a pre-publication methods record.

Sections:
1. Abstract (89 nominees, computational only, no bio claims)
2. Motivation (cationic AMPs, why short variants, why this approach)
3. Seed templates (5 seeds, family rationale, published sources)
4. Candidate generation (3 strategies, conservative groups, rng_seed=2024, 383 total)
5. Scoring pipeline (6 dimensions with formulas and references)
   - Activity-likeness (Scorer 1: heuristic)
   - Boman activity (Scorer 2: Boman 2003 potentials, independent)
   - Model disagreement (|act - boman| uncertainty proxy)
   - Safety proxy (hemolysis risk flags)
   - Synthesis feasibility
   - Novelty (Levenshtein vs 45 curated references)
   - Ensemble (pre-registered weights)
6. Selection criteria (hard filters + ranking + greedy diversity)
7. Results
   - 89 selected from 383 (pass-rate 23%)
   - Per-seed breakdown: SEED-003 best Boman (0.538), SEED-005 best ensemble (0.863)
   - Novelty distribution: 9 high, 58 mid, 22 low
   - Top-10 table with dual-scorer columns
   - Suggested 5-candidate pilot (Boman × low-disagreement criterion)
8. Evidence trail (all artifacts with locations)
9. Reproducibility (make phase3 from clean checkout)
10. What we do not know (mandatory integrity section)
11. Safety and dual-use assessment
12. Next steps (human gates only)
13. References (Boman 2003, Eisenberg 1984, Kyte-Doolittle 1982, Zasloff 1987...)
@cschanhniem
cschanhniem merged commit 9728339 into main Jun 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant